Skip to content

doc: add Helm chart deployment page for Kubernetes - #494

Open
bitflicker64 wants to merge 13 commits into
apache:masterfrom
hugegraph:docs/helm-chart-page
Open

bitflicker64 wants to merge 13 commits into
apache:masterfrom
hugegraph:docs/helm-chart-page

Conversation

@bitflicker64

@bitflicker64 bitflicker64 commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Purpose of the PR

apache/hugegraph#3218 adds a Helm chart for the distributed deployment (PD + Store + Server, optional Hubble), and the site has no pages for it yet. This adds two pages in EN and CN under docs/quickstart/hugegraph/, following the same skeleton, heading numbering and mermaid usage as the sibling pages, both registered in the version route map:

Deployment page (hugegraph-helm, weight 4):

Operations page (hugegraph-helm-operations, weight 5), the chart README's deep-dive content rewritten for an operator: ports and health, scheduling, partition sharding, NetworkPolicy including how NodePort clients appear under kindnet/Calico/Cilium, the safe Store roll procedure (with the GET :8520/v1/partitions follower behavior before and after apache/hugegraph#3232), the disaster recovery runbook (leader targeting, patrol and balance semantics including the apache/hugegraph#3233 refusal body, telling a real run from a no-op, and the empty-PVC Store replacement on both pre- and post-#3234 images), scaling including the Store drain steps, running Hubble outside the cluster, and the two causes of Could not rebind.

The full values reference stays in the chart README; the pages link to its anchors instead of duplicating the tables, so they cannot drift apart.

Checked against the chart at commit 2a6e2c404 of apache/hugegraph#3218 (2026-10-03). Merge order: apache/hugegraph#3218 first; until it lands, the install command and the tree/master/helm/hugegraph links on these pages do not resolve. Both pages are listed in the docs sidebar (data/docs_nav.json) under the HugeGraph components, after PD. Verified locally after merging master (e6389aa), with Python 3.13 and Hugo 0.165.0 Extended: scripts/hugo.sh build, bash dist/validate-links.sh, python -m unittest discover -s scripts -p 'test_*.py' (227 tests, OK), scripts/versioning.py build plus validate for latest (305 pages), and tests/e2e/search-ranking.spec.js against the aggregated site (24 passed). The pages render (mermaid diagram, tables, details blocks, cross-links) in EN and CN.

Planned follow-ups

These are planned work in apache/hugegraph, not conditions for merging this PR. The pages here describe the chart as it is now and get updated when a follow-up changes user-facing behavior.

  • Storage-aware Server readiness. The chart's Server readiness probe calls /versions. Switch it to the storage-aware GET /readiness endpoint once feat(server): add a storage-aware GET /readiness endpoint hugegraph#3221 merges and reaches the images the chart deploys, then update the health-check section here.
  • Upgrade test from a released chart. Upgrades have been tested only between unreleased chart revisions. After the first release that ships the chart, test helm upgrade from that released version to the current chart.
  • Kind install job in CI. helm-chart-ci.yml covers the static layer only (lint, unit tests, render assertions, kubeconform). Add a job that installs values-single.yaml on Kind and runs helm test, with a timeout-minutes budget sized for the Server and Store startup waits, as planned on 2026-08-27 (note on apache/hugegraph#3132).
  • Cross-graph isolation test. From the same note: create two graphs, write to one, and assert the other is unaffected, including after a truncate. It needs no fault injection, so it can run in the Kind job.

Rendered pages

Screenshots show the deployment page; the operations page uses the same skeleton.

EN page: overview with component table, topology diagram, prerequisites and install

CN page: the same skeleton fully translated

Documents the distributed Helm chart from apache/hugegraph#3218:
install with the three values presets, the auth Secrets model,
PD health vs readiness probing, upgrade and scale-down caveats,
uninstall behavior, and current limitations. EN and CN pages under
quickstart/hugegraph, weight 4, matching the sibling page skeleton.
The versioned build validates that every canonical Docs page has a
logical ID in data/version_routes.json; the two new pages were missing
from it, failing the Build latest job. Adds en: and cn: entries with
latest populated and the four older versions null, the same shape as
the other pages introduced after 1.7.

@bitflicker64 bitflicker64 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: yes, until apache/hugegraph#3218 merges. Summary: The EN and CN pages follow the chart at 05f3d9e0b closely and render in CI, but the chart is not on apache/hugegraph master yet (#3132 was closed unmerged, #3218 is an open draft), so merging this first publishes install commands that fail and seven 404 chart links per language. Four smaller mismatches with the chart are inline; each applies to the CN page too. Evidence: compared both pages with helm/hugegraph at 05f3d9e0b (values*.yaml, values.schema.json, _helpers.tpl, NOTES.txt, README anchors); helm template with defaults, values-cluster.yaml, hubble.enabled=true and invalid secret values; gh api repos/apache/hugegraph/contents/helm?ref=master (404); gh pr checks 494 (all passing, publish skipped).


The Helm chart deploys a distributed HugeGraph cluster on Kubernetes: PD, Store, and Server, plus the optional
Hubble UI. It lives in the main repository under
[`helm/hugegraph`](https://github.com/apache/hugegraph/tree/master/helm/hugegraph).

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ important: the PR body already links #3218, so this is only about merge order. The chart is not on apache/hugegraph master yet: gh api repos/apache/hugegraph/contents/helm?ref=master returns 404, #3132 was closed without merging, and #3218 is still an open draft. If this PR merges first, the site publishes a guide where the helm install in 3.2 fails because helm/hugegraph does not exist on master, and all seven tree/master/helm/hugegraph links on the page (EN and CN) are 404s. Please convert this PR to draft until #3218 merges.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Still applies at 795d5f8. Checked 2026-10-02: apache/hugegraph master (176fb56dd) has no helm/ directory, and apache/hugegraph#3218 is open (out of draft, head d05741d9c). Until it merges, the install in 3.2 and the tree/master/helm/hugegraph links fail. Leaving this thread open as the merge-order gate: merge #3218 first, or hold this PR until then.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Still applies at a5fadcc. Checked 2026-10-03: apache/hugegraph master (02628ed50) has no helm/ directory, and apache/hugegraph#3218 is open at 2a6e2c404. Leaving this thread open as the merge-order gate: merge #3218 first, then this PR.

Comment thread content/en/docs/quickstart/hugegraph/hugegraph-helm.md
Comment thread content/en/docs/quickstart/hugegraph/hugegraph-helm.md Outdated
Comment thread content/en/docs/quickstart/hugegraph/hugegraph-helm.md Outdated
Comment thread content/en/docs/quickstart/hugegraph/hugegraph-helm.md Outdated

@imbajin imbajin left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: yes. Summary: The guide contains memory and probe claims that differ from the paired chart and a Store rollout check that can remove a second replica before recovery. Evidence: exact-head comparison with chart #3218 values, helpers, and README; the docs workflow checks passed.

Comment thread content/en/docs/quickstart/hugegraph/hugegraph-helm.md Outdated
Comment thread content/en/docs/quickstart/hugegraph/hugegraph-helm.md Outdated
Comment thread content/en/docs/quickstart/hugegraph/hugegraph-helm.md
Comment thread content/en/docs/quickstart/hugegraph/hugegraph-helm.md Outdated
Comment thread content/en/docs/quickstart/hugegraph/hugegraph-helm.md Outdated
Bring the deployment page level with the chart at 8603cdbb3: the
values-cluster preset now ships NetworkPolicy and 5Gi/8Gi Store memory,
PD startup and liveness derive from the replica count (pd.livenessPath,
single-PD /v1/ready per apache/hugegraph#3222), the Store roll procedure
no longer treats Up in PD as the between-pods check, and the Limitations
follow the merged fixes, including in-place empty-PVC Store recovery on
images carrying apache/hugegraph#3234.

Add an operations page (EN and CN, registered in the version route map)
rewritten for operators from the chart README: ports and health,
scheduling, partition sharding, NetworkPolicy with the per-CNI NodePort
client behavior, safe Store rolls, the disaster recovery runbook with
the post-#3234 procedure, scaling including the Store drain steps,
running Hubble outside the cluster, and the two causes of "Could not
rebind".

Verified: scripts/hugo.sh build passes and both new pages render with
their cross-links in EN and CN.
bitflicker64 added a commit to hugegraph/hugegraph that referenced this pull request Sep 24, 2026
Cut the README from 1,472 to 1,003 lines now that the docs site carries
the operator walkthroughs. Each moved section keeps its load-bearing
warning and commands and links its docs-site path: NetworkPolicy details,
Cluster Health, Scheduling, Partition Sharding, the Disaster Recovery
narrative (the keep-the-PVC rule and the retirement commands stay), the
Scaling procedures, the outside-Hubble paths, and the Could-not-rebind
measurements. The quickstart, presets, Kind flow, upgrade warnings with
the OnDelete Store roll, the values tables, the validation list, every
troubleshooting symptom and check command, and the Limitations stay.

Deduplicate repeated passages to one home each: the ordinal-truncation
explanation (Release Name Too Long), the anti-affinity trade
(Installing), and the Hubble single-replica/H2 constraints (the Hubble
section; the values.yaml comment now points there).

README-only plus a values.yaml comment: helm template output is
byte-identical, lint passes on the three presets, and all 139 unit tests
pass. The docs-site links resolve once apache/hugegraph-doc#494 merges.
@imbajin
imbajin requested a lite review from Copilot September 25, 2026 08:07

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Several documentation corrections and navigation updates remain unresolved.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 6 Medium severity

Open (6)
What changed in this PR

Adds bilingual Helm deployment and operations documentation for HugeGraph on Kubernetes.

Changes:

  • Adds English and Chinese deployment guides.
  • Adds operations guides covering installation, networking, scaling, recovery, and Hubble.
  • Registers the new pages in the version route map.
File Description
data/​version_routes.json Adds routes for the new pages.
content/​en/​docs/​quickstart/​hugegraph/​hugegraph-helm.md English deployment guide.
content/​en/​docs/​quickstart/​hugegraph/​hugegraph-helm-operations.md English operations guide.
content/​cn/​docs/​quickstart/​hugegraph/​hugegraph-helm.md Chinese deployment guide.
content/​cn/​docs/​quickstart/​hugegraph/​hugegraph-helm-operations.md Chinese operations guide.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread content/cn/docs/quickstart/hugegraph/hugegraph-helm-operations.md
Comment thread content/cn/docs/quickstart/hugegraph/hugegraph-helm.md Outdated
Comment thread content/cn/docs/quickstart/hugegraph/hugegraph-helm.md Outdated
Comment thread content/en/docs/quickstart/hugegraph/hugegraph-helm-operations.md
Comment thread content/en/docs/quickstart/hugegraph/hugegraph-helm.md Outdated
Comment thread content/en/docs/quickstart/hugegraph/hugegraph-helm.md Outdated

@imbajin imbajin left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: yes. Summary: The recovery guide compares a term that this Store endpoint does not return, can run balancing commands during a patrol-triggered cooldown, and leaves a re-used ordinal with a Tombstone Store ID on its retained PVC. Evidence: current Server master handlers and the chart head's default PVC-retention setting.

Comment thread content/en/docs/quickstart/hugegraph/hugegraph-helm-operations.md Outdated
Comment thread content/en/docs/quickstart/hugegraph/hugegraph-helm-operations.md
Comment thread content/en/docs/quickstart/hugegraph/hugegraph-helm-operations.md Outdated
Comment thread content/en/docs/quickstart/hugegraph/hugegraph-helm-operations.md Outdated

@bitflicker64 bitflicker64 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: yes. Summary: The deployment page's Upgrade section says the first upgrade after install rolls PD and Server together but never warns that a Server started during that PD roll can lose its Gremlin binding for the life of the Pod, and its scaling paragraph says only PD and Store shrinks are rejected although the chart also rejects PD growth. Evidence: compared both EN and CN pages with helm/hugegraph at 3996bb12 (head of apache/hugegraph#3218): README.md Upgrading the Chart, templates/_helpers.tpl lines 649-681, values*.yaml, NOTES.txt, networkpolicy.yaml, README anchors; findings already raised in earlier reviews are not repeated.

Comment thread content/en/docs/quickstart/hugegraph/hugegraph-helm.md Outdated
Comment thread content/en/docs/quickstart/hugegraph/hugegraph-helm.md Outdated

@imbajin imbajin left a comment •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Score: 8.5/10

Before merging:

  • Coordinate with Helm #3218 and settle the operational wording.
  • Resolve the missing scripts/oink_module.py prerequisite and rerun website checks.
  • Keep artifact/version-specific recovery limitations explicit.

The reported site failure occurs before it can establish a Helm documentation regression.

Describe the chart as of apache/hugegraph#3218 at 0e754b98c, in EN and CN:

- helm test now queries Gremlin on every Ready Server Pod through the
  headless Service, with the replica floor and the 150 s retry window
- values-cluster.yaml sets Store updateStrategy OnDelete; the other
  presets keep RollingUpdate, and OnDelete is not a safety check by itself
- rollingUpdate.maxUnavailable accepts only the integer 1
- Server startup budget: guaranteed time is (failureThreshold - 1) *
  periodSeconds, default 91 x 5, HG_SERVER_STARTUP_TIMEOUT_S 150 s
- PD secret and admin password whitespace and backslash constraints, and
  the Hubble egress render guard with server.advertiseUrl
- Limitations and troubleshooting for apache/hugegraph#3226 and #3228,
  #3229 referenced from the Upgrade section, and the Scaling link moved
  from the chart README to the operations page where the procedure lives
Since the OINK migration the sidebar comes from data/docs_nav.json, and
the two Helm pages were not in it, so no navigation led to them. Add both
under the HugeGraph components entry after PD, regenerate the derived maps
with materialize_docs_navigation, and update the pinned navigation stats:
latest gains two pages, and each archived version, which lacks the pages,
counts two more removed entries.
State each limitation as the behavior a user sees, with the issue link
kept as a reference. The Store roll now says plainly that no endpoint
reports restore completion. The uninitialized-PD entry separates what was
observed outside Kubernetes from what the chart's probes do, and no longer
offers an untested recovery step. helm test prints the failing Pod's IP,
not its name.
kubectl port-forward stays in the foreground, so the curl commands that
followed it in the same block never ran. Split the install check, the
operations health check and the per-Pod Gremlin check into a forward
block and a call block, in English and Chinese.
- values-cluster.yaml adds the Server PodDisruptionBudget; values.yaml
  already enables the PD and Store budgets.
- Each JVM sets its heap to half of the free node memory, so the
  single-node preset needs memory sizing too.
- Explain that --reuse-values keeps old values as the base and how to
  adopt new defaults; Prerequisites referred to this.
- The chart rejects PD growth as well as PD and Store shrinks.
- The Store per-group partition route never fills term.
- patrolPartitions also sets the 180 s balance-shard flag when it
  reallocates a shard group.
- PD exposes POST /v1/members/change behind its REST authentication.
- A retained PVC of a retired Store ordinal cannot register again.
The admin password must now be printable ASCII with no spaces, colons
or backslashes, and a bring-your-own Secret is checked at Pod start.
A no-change upgrade rolls nothing, and a template-only pipeline rolls
the Pods that read a generated credential on every sync instead of
leaving the annotation inert.

Document the Server and Hubble Service exposure acknowledgements, the
install-time identity and bootstrap-only values, and why helm rollback
bypasses the chart's guards. EN and CN updated together.

@bitflicker64 bitflicker64 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: no. Summary: Both EN and CN pages match the paired chart at apache/hugegraph#3218 head 2a6e2c404 on every value, guard, probe path, Secret rule and README anchor checked, and the earlier inline threads are addressed at this head. One new Minor point is inline; the open merge-order thread on line 15 of the EN deployment page still holds (apache/hugegraph master has no helm/ directory and #3218 is open), so this PR should merge only after #3218. Evidence: compared both pages with helm/hugegraph at 2a6e2c404 (values.yaml, values-single.yaml, values-cluster.yaml, values.schema.json, _helpers.tpl, networkpolicy.yaml, tests/test-connection.yaml, hubble-pvc.yaml, README headings); PD StoreAPI /shardGroups response key; gh api repos/apache/hugegraph/contents/helm?ref=master (404); gh pr checks 494 (all passing, publish skipped).

helm uninstall hugegraph --namespace hugegraph
```

Two kinds of state survive on purpose. PersistentVolumeClaims created by the StatefulSets are kept (Kubernetes

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor: this section lists two kinds of state that survive uninstall, but there is a third. With hubble.enabled=true and hubble.persistence.enabled=true, templates/hubble-pvc.yaml at chart 2a6e2c404 annotates the Hubble PVC with helm.sh/resource-policy: keep, and its comment says the volume holds the stored connection metadata and credentials. An operator who follows this page and deletes only the StatefulSet PVCs leaves that volume, with its credentials, in the namespace. Please add the Hubble PVC here (for example: "With Hubble persistence on, its PVC is kept too; delete it to discard Hubble's stored connections and credentials") and make the same change in CN section 8, line 261.

@imbajin imbajin left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: yes. Summary: The Store runbook overstates restart and disruption protection and can treat an empty-PVC replacement as recovered before replication catches up; the external Hubble instructions omit authenticated PD credentials and transport safeguards. Evidence: chart #3218 at 2a6e2c404; TaskScheduleService and PartitionEngine source; Kubernetes disruptions documentation.

All ports are configurable through values; changing `server.port` updates the
listener, container port, and Service together. A stalled component (process
alive but frozen) is ended by its liveness probe: the default 20 s period and
3-failure threshold bound the blast radius of a stalled Store at roughly one

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Important: The 20 s period and three failures describe detection in about one minute, but the chart sets store.terminationGracePeriodSeconds to 300 s. Kubelet can wait through that grace period before forcing down an unresponsive Store, so the total restart and recovery window can be several minutes longer. Please call one minute the approximate detection time and mention the grace period; sync the Chinese page.

and `peers` null for followed groups. The per-group path works on both.

Leave a margin after the membership check, keep `store.pdb.minAvailable`
at `replicas - 1` so an accidental second eviction is refused, and treat a

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Important: This procedure tells operators to delete Store Pods one at a time, then says the PDB refuses a second eviction. If operators use ordinary Pod deletion, Kubernetes bypasses PDB enforcement; only Eviction API requests honor the budget. The PDB therefore cannot protect the manual rollout as written. Please specify an eviction-based procedure or limit this claim to kubectl drain and other eviction clients; sync Chinese. Kubernetes documentation.

one), `POST /v1/store/<oldId>` with `{"storeState":"Tombstone"}` on the PD
leader, run `GET /v1/task/patrolPartitions`, and wait; verify every shard
group is back to full shard count with one leader, no group names the old
ID, and the replaced Store's own `:8520/v1/partition/<groupId>` answers

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Important: Full shard-group membership and HTTP 200 from this endpoint do not prove the replacement has caught up. PartitionEngine.syncShardIdentities() updates the shard IDs in PD before waitForReplicate() completes, and this endpoint exposes the current logIndex during that wait. Please compare the replacement's index with the leader or a current peer for every group before deleting the old Tombstone record; sync Chinese.


**PD discovery** (the outside Hubble asks PD for the Server address):
in-cluster names such as `*.svc` are not reachable from outside, so the
chart provides two knobs. Set `server.advertiseUrl` to the absolute

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Important: PD returns this advertised URL to Hubble, which then sends Server login credentials to the discovered URL. This section allows http:// without the HTTPS or trusted-channel warning given for server.direct_url above. Please require HTTPS or a trusted private transport for production advertised URLs and sync Chinese.

with:

```properties
pd.enabled=true

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Important: This PD-mode example omits Hubble's PD REST credentials. operations.pd.password defaults to empty, while the chart's default PD REST authentication requires credentials; chart-managed Hubble injects operations.pd.username=hubble and the PD secret. As written, the external Hubble operations collector gets authentication failures. Add these properties to the example and show how to provide the PD credential securely; sync Chinese.

- clarify readiness and Server hook acceptance
- document coordinated credential maintenance
- require separate Hubble login and query checks
- align recovery guidance in both languages

@imbajin imbajin left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: yes. Summary: The rolling Store gate can pass while a Learner has not been promoted, which can leave a shard without quorum. The inline comments also cover policy, credential-rotation and PD rebalance guidance. Evidence: Verified source locations are linked in the inline comments.


Delete the next Store only when the replaced Pod is `Ready`, its Store id
shows a fresh `lastHeartBeat` in `/v1/stores`, and every group reports its
full shard count with exactly one `Leader`.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

‼️ critical. Blocking: yes. Summary: This gate counts every shard entry, so a Leader+Follower+Learner group satisfies the full-count/one-Leader condition before the new replica is promoted; deleting another Store can leave only one voter. Please require one Leader and only Followers (no Learner/None) before continuing, including the shrink gate, and sync the Chinese instructions. Evidence: PartitionEngine adds peers as Learners and promotes them only after waitForReplicate; PD treats Learner as migration in progress.

namespace and `curl` the PD client Service on the REST port: it must time
out.

With the policies on, the release admits only its own traffic:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ important. Blocking: no. Summary: The statement that the release admits only its own traffic holds only when no other applicable NetworkPolicy allows traffic to these Pods; Kubernetes unions allowed rules, so a broader policy can admit sources omitted from extraIngress. Please qualify this boundary and advise operators to audit other policies selecting these Pods in both languages. Evidence: Kubernetes documents additive NetworkPolicy rules and the union of allowed connections.

<summary>Rotation and caveats</summary>

- The admin password is applied only when the auth metadata is first created, so changing the Secret later does not
rotate an existing cluster's password. Rotate it through the Server's auth API instead.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ important. Blocking: no. Summary: Rotating only through the Server auth API leaves the chart admin Secret unchanged, so the documented password lookup and helm test hook keep using the old credential. Please set the Secret to the API value and roll Server replicas to refresh their auth caches, or link to the chart README rotation procedure; sync Chinese. Evidence: The test hook reads PASSWORD from the chart Secret; the chart README requires Secret synchronization and a Server rollout.

[apache/hugegraph#3231](https://github.com/apache/hugegraph/issues/3231).

Run `patrolPartitions` after replacing a Store that is not coming back,
`balancePartitions` once the cluster is stable again, and `balanceLeaders`

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ important. Blocking: no. Summary: Tombstoning a Store starts a separate 30-minute cooldown: balancePartitionShard() returns null during that interval, even after the documented 180-second patrol window. Please document the Tombstone cooldown before advising balancePartitions, and distinguish it from the BALANCE_SHARD_KEY timer. Evidence: The status listener records Tombstone time; the cooldown constant and early return and the 30-minute guard.

@bitflicker64 bitflicker64 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: yes. Summary: Both pages are accurate against the chart merged in apache/hugegraph#3218 (0a3e4ae5f), and the EN and CN pages match. The Store shrink procedure in operations section 8 fails at its Tombstone step on chart-deployed clusters, and Store growth is not documented; neither point is in an earlier thread, and the open threads are not repeated here. Evidence: helm template (helm 4.2.4) of helm/hugegraph at apache/hugegraph master 8fd9841eb with the default, values-single, values-cluster and hubble.enabled=true values, and with store.replicas 3 against 5; the PD source (StoreNodeService.updateStore, StoreAPI.setStore, PDRestService.updateStore, PDConfig, TaskScheduleService.init) at the same commit; the README anchors and the EN/CN code-span diff; gh pr checks 494, all passing with publish skipped.

`store.replicas` is at least 3).
2. Map the ordinals the shrink will delete (the highest ones) to Store
ids through `/v1/stores`, matching on the Pod address.
3. `POST /v1/store/<id>` with `{"storeState":"Tombstone"}` for each

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Important: Step 3 cannot succeed on a cluster this chart deployed. The chart sets HG_PD_INITIAL_STORE_COUNT to store.replicas (pd-statefulset.yaml L123-124), and that value is PD's minStoreCount (PDConfig L63-64). POST /v1/store/<id> goes through PDRestService.updateStore into StoreNodeService.updateStore, which refuses a Tombstone for a Store whose last state is Up when getStores().size() - 1 < minStoreCount (StoreNodeService L271-278). On a fresh 5-Store install, retiring a healthy Store gives 5 - 1 < 5, so PD returns LESS_ACTIVE_STORE ("The number of active stores is less then 5"). StoreAPI.setStore serializes that exception into the JSON body instead of failing the request, so an operator following the steps may not notice. Step 4 then never finishes. The Disaster Recovery retirement passes this check only because the replacement adds a second record at the same address. I traced this in the source at apache/hugegraph master 8fd9841eb and did not run it on a cluster. Please say that PD refuses to Tombstone an Up Store while the chart pins pd.initial-store-count to the current store.replicas. Then either document a path that works or mark Store shrink as unsupported for now, as the page already does for PD. Please make the same change in the Chinese page, section 8 step 3 (line 297).


Changing PD or Store replicas on a live release is not a values change:
raft and shard membership are persisted, and Pods alone do not reconfigure
them. The chart rejects both directions for PD and a shrink for Store by

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor: This section and the deployment page treat Store growth as a plain values change, but three things about it go undocumented. (1) With the default store.pdb.minAvailable: 2, --set store.replicas=5 fails to render with "store.pdb.minAvailable must be at least store.replicas - 1" (_helpers.tpl L780-781); store.pdb.minAvailable has to move with the replica count in both directions, which also affects the helm upgrade in step 5 of the shrink below. (2) HG_PD_INITIAL_STORE_LIST and HG_PD_INITIAL_STORE_COUNT live in the PD Pod template, so growing Store rolls every PD Pod. A helm template diff of store.replicas 3 against 5 shows only those two PD env values changing. (3) PD schedules no balance task (TaskScheduleService.init), so new Stores hold no partitions until balancePartitions runs on the PD leader. Please add a short "Store, growing" paragraph that covers these three points and then the check from section 6 (full shard count, one Leader). Please sync the Chinese page, section 8 (line 283).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants